Algorithms for Molecular Biology
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match Algorithms for Molecular Biology's content profile, based on 17 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Bohnenkaemper, L.; Stoye, J.
Show abstract
The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.
Subramanian, G.; Thiel, W.; Singh, R.
Show abstract
Aptamers are structured nucleic acid ligands capable of high affinity, high specificity molecular recognition generated using variations of the SELEX (Systematic Evolution of Ligands by Exponential Enrichment) process. However, SELEX often produces sequences that enrich yet may lack binding efficacy. We propose a measure called the Ruggedness Composite Index (RCI) along with a method for computing it, that can be used to distinguish binding-competent ('active') aptamers from weak or non-binding ('inactive') aptamers. Given a set of aptamers, RCI incorporates information on their fragmentation (landscape partitioning), basin entropy (metastable state distribution), cumulative density irregularity (non-uniform occupancy), and structural energy correlation length (structure-energy coupling scale). We test whether secondary-structure folding energy landscape topology distinguishes active from inactive aptamers using a multiscale level set framework across six datasets. Active aptamers show lower RCI values and occupy smoother, funnel-like conformational spaces, while inactive aptamers show higher RCI values, reflecting fragmented, high-entropy landscapes. By contrast, classical thermodynamic features, such as minimum free energy, show limited discrimination between active and inactive aptamers. In all datasets, sequences that exhibit enrichment which is not monotonic but lack specificity exhibit elevated ruggedness, indicating landscape topology can predict non-specific enrichment. These results indicate that folding landscape organization can be used as a predictor of aptamer activity and establish RCI as a simple, mechanistically interpretable measure for improving candidate prioritization, especially in therapeutic aptamer discovery.
Seiler, E.; Willemsen, M.; Piro, V. C.; Reinert, K.
Show abstract
Motivation: A continued decrease in sequencing costs has facilitated the exponential increase in available sequencing data, with public databases like the European Nucleotide Archive (ENA) and Sequence Read Archive (SRA) reaching well in the order of petabases. This has been the incentive to develop more scalable tools for common bioinformatics tasks. One such task is the approximate searching of short sequence patterns like genes or reads in reference data sets. In recent years, a variety of indexing data structures have been proposed for searching large sequencing databases. The state-of-the-art index, the Hierarchical Interleaved Bloom Filter (HIBF) was first-in-class to index one million samples. To be useful for expanding repositories, it must be extended to support dynamic updates. Results: In this paper, we introduce a scalable and updatable sequence-search index by extending the HIBF with partial rebuilding to support efficient updates. We demonstrate the Dynamic HIBF's capacity for large-scale data by iteratively creating an index from over 100 TB of compressed reads across more than 39,000 full human RNA-Seq samples, updated in consecutive batches of 100. To benchmark against state-of-the-art tools, we evaluated incremental performance on a subset of 5,000 samples sub-sampled to 1% of their original read depth. In this comparative setting, the dynamic HIBF completed the sequential insertion of all 5,000 samples within 5 hours--24 to 65 times faster than competing methods and twice as fast as the static HIBF.
Krieg, R.; Becker, F.; Saenko, S.; Diehl, J.; Stanke, M.
Show abstract
Scaling the structural annotation of protein-coding genes to all eukaryotic genomes remains a major challenge. While recent deep learning methods rival evidence-based pipelines without requiring RNA-seq or alignments, they are entirely supervised. They depend on large, high-quality training sets from diverse genomes, leaving many basal eukaryotic clades without an accurate ab initio gene finder. We present Vipsania, the first unsupervised deep gene finder. A differentiable hidden Markov layer inside a deep sequence model learns to predict gene structures from unannotated genomes alone. Vipsania is pretrained for virtually all eukaryotes and finetunes without supervision on the target genome. It is, on average, more accurate than supervised methods across most clades and avoids the accuracy drop that supervised models suffer on distant target genomes. Vipsania adapts to non-standard genetic codes and provides a fast and highly versatile tool for unbiased, pan-eukaryotic genome annotation. The source code is available at https://github.com/gaius-augustus/vipsania.
Zhao, C.; Ji, Z.
Show abstract
Spatial domain detection is a central task in spatial transcriptomics, yet existing methods exhibit highly variable performance across datasets. We introduce L-STAR, a visual LLM-guided, consensus-based framework that leverages the visual reasoning capacity of large language models to adaptively rank and integrate spatial domain detection methods. L-STAR achieves robust and consistently improved performance, outperforming single spatial domain detection methods across diverse datasets.
Siemers, M.; Lopez, J. L.; Dutilh, B. E.
Show abstract
Bacteriophages can only be understood through their interactions with bacterial hosts. As environmental sequencing efforts expanded, the number of available phage genome sequences has exploded, yet the vast majority of these sequences lack host information. Predicting the host of a newly observed phage is therefore a key challenge in virology. Several computational tools can predict phage-host relationships from genomic data, but they share notable limitations: (1) the number of different hosts that can be predicted remains relatively restricted; (2) tools tend to assign confident host predictions to non-viral input sequences; and (3) most tools have a trade-off between accuracy and speed. Here we present PhageTransformer (PT), a deep learning model for phage-host prediction that addresses these limitations. We benchmark PT against existing tools on 3,881 independent phage-host pairs from GenBank and public HiC data, and demonstrate that it achieves competitive or superior prediction accuracy at greatly reduced runtime.
Zeng, Z.; Wang, Y.
Show abstract
Motivation: The Interactive Tree of Life (iTOL) is widely used to display and annotate phylogenetic trees, but managing its format-sensitive annotation files impede reproducible high-throughput analyses. Among the maintained Python packages and versions evaluated, none combined template generation, taxonomic monophyly assessment and iTOL batch operations. Results: PyiTOL validates inputs, generates 31 iTOL template schemas (22 accepted by the live batch uploader), performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay. On a topology-constructed benchmark, all calls matched prespecified labels for 4,389 groups; on a 700-genome tree, binary mono/non-mono calls agreed with ETE4 for 409 genera; 17,294 GTDB R232 genera were processed in about 17 s. Availability and Implementation: PyiTOL 1.0.3 (Python [≥]3.10; Linux, macOS and Windows) is MIT-licensed at https://github.com/ZengZichao/PyiTOL and archived with test data at Zenodo (https://doi.org/10.5281/zenodo.22106806).
Bui, T.-C.; Lee, J.; Ko, J.
Show abstract
Scoring biomolecular complexes is central to structure assessment and drug discovery, yet the complexes themselves vary widely in pose, size, and molecular composition. A scoring function tuned for one interaction type rarely carries over to another, and most existing methods compound the problem by leaning heavily on task-specific labels. We introduce OmniScore, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads. OmniScore couples a graph view and a sequence view of each structure, encodes its three-dimensional geometry, and compresses representations into a compact latent space that a reconstruction module and prediction heads can reuse. We pretrain this backbone on diverse datasets including complexes, monomers, and small molecules with complementary objectives: coordinate recovery, correcting corrupted input tokens, predicting molecular identity, and grounding the representation in structure-level physical quantities. Across the evaluated benchmarks, OmniScore gave the best antibody-antigen and nanobody-antigen quality assessment on all reported metrics compared to state-of-the-art baselines. Its frozen residue embeddings matched the state-of-the-art protein-tokenization method with an average functional-site accuracy of 71.8% on a standard residue-level benchmark. On protein-ligand scoring and ranking benchmarks, it performed on par with methods built specifically for that single task. These results suggest that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend on interfacial and residue-level structure, within the evaluated settings.
Xuan, H.; Huang, Y.; Bian, J.; Liu, X.
Show abstract
Motivation: Interactive tools that let non-programmers explore an analyzed single-cell dataset, its embeddings, gene expression, cell metadata, and marker genes, have become standard laboratory infrastructure. Every actively maintained tool in this space (ShinyCell, ScRDAVis, sCIRCLE, scViewer) is built on R Shiny and requires a Seurat object as input. Laboratories whose primary analysis pipeline is Python/scanpy, the dominant framework for single-cell RNA-seq, spatial, and multi-omic analysis, therefore have no lightweight, language-native option that pairs a shareable web-based viewer with a scriptable Python API: sharing a scanpy result means either exporting to Seurat first or handing over a notebook that only a programmer can run. Results: We present scPyviewer, a web-based viewer that ingests AnnData objects directly and reproduces the core interaction patterns of the incumbent R Shiny tools without leaving the Python stack. In a feature-parity audit against three actively maintained R Shiny incumbents, scPyviewer matches or exceeds every baseline capability (7/7); among these, it uniquely offers native AnnData ingestion with no Seurat conversion, and cross-dataset comparison over shared genes and matched cell-type composition. Benchmarked head-to-head against the R/Seurat rendering substrate the incumbents are built on, identical operations, identical data, across three datasets spanning 22,315 to roughly 313,000 cells, scPyviewer renders every core view faster at every scale tested (up to 3.6x on a single view) and at a fraction of the memory (5.2x lower on the smallest dataset). At the largest scale tested, the gap becomes categorical rather than incremental: scPyviewer completes every view on a 313,000-cell dataset while the Seurat substrate exhausts an 8 GB memory budget and fails outright. Beyond the interactive app, scPyviewer installs via pip or conda and exposes a public Python API that returns Matplotlib figures and pandas tables for scripted, publication-ready output. Availability and implementation: scPyviewer is implemented in Python 3.11 (scanpy 1.11.5, anndata 0.12.19, streamlit 1.59.2, plotly 6.9.0) and distributed with a one-command reproduction interface that installs pinned dependencies, regenerates the benchmark and all figures, and launches the interactive app. Source code is available at https://github.com/xuan13hao/scPyviewer.git.
da Silva, K.; Sarkodie, S.; Marques, K.; Vieira, P.; Oliveira, R. D. d.; Pereira dos Santos, P. C.; Moreira Puga, M. A.; Costa, A. G.; Gregorio Machado, J. P.; Spener-Gomes, R.; Yang, E.; Savic, R.; Cordeiro-Santos, M.; Croda, J.; Andrews, J. R.
Show abstract
Background: Polymorphisms in the N-acetyltransferase 2 (NAT2) gene explain much of the interindividual variation in isoniazid (INH) metabolism and determine risk of toxicities. However, there is limited evidence to guide INH dose adjustment according to the NAT2 acetylator profile in weekly rifapentine-INH tuberculosis preventive therapy (TPT). Methods: In a prospective, multicenter, within-subject PK trial (NCT05413551), adults initiating 3HP in Brazil were assigned genotype-guided INH doses (slow: 5 mg/kg <=300 mg; intermediate: 15 mg/kg <=900 mg; rapid: 25 mg/kg <=1,500 mg) alongside a standard 900 mg flat dose on an alternate occasion. AUC0-24 and C24 were estimated from serial blood samples; a two-compartment Michaelis-Menten population PK model characterized NAT2 effects on clearance. Results: Among 228 participants, 47.4% (108/228) were intermediate, 43.4% (99/228) slow, and 9.2% (21/228) rapid acetylators. Genotype-guided dosing reduced AUC0-24 variability approximately two-fold versus standard dosing (CV 58.8% vs 76.8%) and increased exposure uniformity (median AUC0-24 27.2 [IQR 18.8-41.3] vs 43.2 [27.3-71.0] mg h/L). Among slow acetylators, C24 >0.15 ug/mL decreased from 27/42 (64%) with standard dosing to 1/42 (2%) with genotype-guided dosing (P<0.0001). In 104 participants with intensive PK sampling, rapid acetylators receiving guided doses had AUC0-24 similar to standard-dose intermediate acetylators (42.8 vs 39.5 mg h/L; P=.63). Monte Carlo simulations supported doses of 600, 900, and 1,200 mg for slow, intermediate, and rapid acetylators, respectively. Conclusions: NAT2-guided isoniazid dosing reduced variation in drug levels, averting very low and high AUC and C24. These findings inform genotype-stratified dosing of INH for TPT, which might reduce toxicities and improve outcomes.
Zink, T.; Noren, H.; Valdivia, D.; Yohn, C.; Hundal, J.; Chen, S.; Scarisbrick, D.; Sun, H.
Show abstract
Abstract: Objective: Post-traumatic epilepsy (PTE) is a common sequela of traumatic brain injury (TBI). Research indicates that individuals with PTE tend to experience greater cognitive difficulties compared to those with TBI alone. However, it is plausible that a distinct cognitive profile exists that distinguishes between TBI cases with and without PTE. We aimed to identify longitudinal changes in cognitive measures among TBI patients to better assess the changes associated with developing PTE. Setting: Outpatient. Participants: Prospective subjects who had suffered TBI within 6 months post-injury (TBI-6M, n=32), retrospective subjects with pre-existing PTE diagnoses (PTE, n=20), and healthy control subjects (HC, n=41). Design: We examined cognitive performance for TBI patients within 6 months post-injury, then again within 12 months (TBI-12M, n=26), and within 18-months (TBI-18M, n=25), and compared this with cognitive performance among HC and PTE. Main Measures: Cognitive tests administered yielded 15 test components for analysis. We utilized linear mixed effects modeling to examine cohort-level differences cognitive function. Results: 11/15 tests showed a significant performance deficit in the PTE subjects compared to HC. TBI-6M was not significantly different from the PTE subjects; with time, 9/15 tests showed some degree of recovery in TBI subjects. Tests for information processing speed/working memory and executive function showed strong recovery (TBI-6M vs. TBI-18M, SDMT written: p<0.0001, SDMT oral and COWAT: p<0.001). Tests for visual attention/working memory also showed a smaller but significant recovery (TBI-18M vs. PTE, p<0.05). By contrast, tests for verbal memory [HVLT-R Delayed Recall] showed chronic impairment in TBI (TBI-18M vs HC, p<0.0001). TBI subjects generally trend towards recovery in cognitive performance post-TBI. Conclusions: Information processing speed/working memory are strong indicators for TBI recovery, while auditory learning/memory shows chronic impairment. The stagnation of recovery in cognitive domains typically characterized by robust recovery may correlate with an elevated risk of developing PTE.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.
Show abstract
Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.
Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.
Show abstract
Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.
Langbaum, J. B.; Erickson, C. M.; Langlois, C.; Wood, E. M.; Egleston, B. L.; Harkins, K.; Mim, R.; John, S.; Brown, C.; Brown, S.; Howe, S.; Cacioppo, C.; Eppelmann, L.; Enos, J.; Salata, H.; DeSantiago, D.; Largent, E. A.; Reiman, E. M.; Denkinger, M. N.; Ashton, N. J.; Roberts, J. S.; Karlawish, J.; Bradbury, A. R.
Show abstract
Importance: Patients are increasingly learning Alzheimers disease (AD) genetic and biomarker results through electronic health portals. Evaluation of alternative scalable delivery models for return of AD risk information is needed to best support patient understanding and psychological well-being. Objective: To determine whether a patient-centered digital platform is comparable to clinician-mediated telehealth sessions for returning APOE and plasma pTau-217 results on outcomes of knowledge and psychological well-being. Design: The Evaluation of Self-Mediated Alternatives for Risk Testing Education and Return of Results (eSMARTER) study was a noninferiority trial of a patient-centered digital platform compared to clinician-mediated disclosure of APOE genotype and optional pTau-217 disclosure. Setting: Decentralized, fully remote trial enrolled participants in the contiguous United States (U.S.) between October 2024 and February 2025, with follow-up completed in November 2025. Participants: Eligible participants were aged 60-80 and had previously undergone APOE genotyping (without disclosure) via the GeneMatch program, passed psychological screening, had internet access, and were English-speaking. Interventions: Participants were randomized, 2:1, to the eSMARTER digital platform or clinician-mediated disclosure of APOE genotype. Following the 6-month post-APOE assessment, participants were offered optional pTau-217 disclosure via the same randomized modality. Main Outcomes and Measures: Primary outcomes at 1-7 days following APOE disclosure included changes in anxiety, disease-specific distress, and AD-related knowledge within a priori non-inferiority margins. Results: 674 persons (mean [SD] age 68 [4.7] years; 451 [67%] female; mean [SD] telephone MoCA=19 [2]) were eligible and provided demographic information. 651 participants were randomized to clinician-mediated (n=216) or digital disclosure (n=435) and completed APOE disclosure (66 [10%] APOE4 homozygotes, 377 [58%] heterozygotes, 208 [32%] non-carriers). 604 participants completed the study; 500 completed optional pTau-217 disclosure. Baseline characteristics were balanced across groups. At 1-7 days following APOE disclosure, scores on AD-related knowledge, PROMIS Anxiety, and disease-specific distress measures met non-inferiority. Conclusions and Relevance: Disclosure of APOE genotype by the eSMARTER digital platform is non-inferior to clinician-mediated telehealth disclosure. No significant between group differences were found following disclosure of pTau-217 results. Together, these results suggest that this digital platform may provide an evidence-based scalable approach for returning AD genetic and biomarker results.
Rohd, S. B.; Thorup, A. A.; Wilms, M.; Schiavon, M.; Streyma, D. H. B.; Laursen, A. F.; Bundgaard, A. F.; Sondergaard, A.; Krantz, M. F.; Veddum, L.; Hjorthoj, C.; Greve, A.; Mors, O.; Nordentoft, M.; Hemager, N.; Gregersen, M.
Show abstract
Objective: This study examined the prevalence of psychotic experiences (PE) and how early onset and persistence of PE contribute to risk and severity of mental disorders in adolescents at familial high-risk of schizophrenia (FHR-SZ) or bipolar disorder (FHR-BP) and adolescents from a population-based control group (PBC). Methods: This is the second follow-up of a nationwide cohort study including 522 children at FHR-SZ (N=202), FHR-BP (N=120), and PBC (N=200). Participants were assessed at ages 7, 11, and 15 using a semi-structured interview to evaluate PE and mental disorders. Results: At age 15, adolescents at FHR-SZ reported more PE than PBC over the past six months (current) and the past four years, while adolescents at FHR-BP only reported more current PE. PE reported at two or three timepoints (persistent PE) predicted any Axis I disorder in mid-adolescence, corresponding to three- (OR 2.9, 95% CI [1.5-5.7]) and 21-fold (OR 21.4, 95% CI [2.8-162.3]) increased risks, respectively. Persistent PE also predicted multimorbidity, with three- (OR 2.8, 95% CI [1.0-7.6]) and four-fold (OR 4.1, 95% CI [1.2-14.1]) increased risks, respectively. This was after adjustment for sex, early mental disorders, and familial risk. Conclusions: This study demonstrates a strong link between persistent PE and mid-adolescence mental disorders. Our findings emphasize PE as important risk markers for mental disorders during mid-adolescence and highlight the importance of monitoring children with PE before age 7 who develop persistent symptoms.
Ndiaye, A.; Thiebaut, A. C. M.; Borel, P.; Sabran, C.; Elis, S.; Guerif, F.; Maillard, V.
Show abstract
The distribution of fat-soluble compounds (including antioxidants) in follicular fluid (FF) remains sparsely documented in relation to in vitro fertilization (IVF) outcomes and existing studies have reported diverging associations. This study aimed to describe plasma and FF concentrations of fat-soluble micronutrients in women undergoing IVF and to analyze their adjusted associations with ovarian function, embryo development and pregnancy outcomes. In 2021-2022, plasma and FF samples were collected from 82 women (first IVF cycle) at oocyte puncture, along with lifestyle data covering the three preceding months. Eleven compounds (two tocopherols, three xanthophylls, five carotenes and retinol) were quantified. All compounds were detected in both compartments (lowest in FF) except phytoene, undetectable in FF. Plasma and FF -tocopherol concentrations were positively associated with plasma estradiol levels before oocyte puncture (both p<0.01) while FF -carotene and lycopene were inversely associated with plasma progesterone concentrations (p=0.01 and 0.02, respectively). Plasma phytofluene and phytoene were positively associated with mature oocyte rate (p=0.03 and p=0.01, respectively), while FF retinol was negatively associated (p=0.03). Carotenes, tocopherols and retinol were inversely associated with later IVF outcomes: fertilization rate (p<0.001 for plasma g-tocopherol, 0.02 for FF retinol), top-quality embryo (p=0.02 for plasma phytofluene), biochemical pregnancy at day 7 post-embryo transfer (p=0.05 for plasma -tocopherol, 0.02 for plasma -carotene), clinical pregnancy (p=0.03 for plasma -tocopherol, 0.01 for plasma phytoene) and live birth (p=0.04 for plasma -tocopherol, 0.02 for plasma phytoene). Plasma and FF g-tocopherol were positively associated with embryo fragmentation (both p<0.05). Finally, among xanthophylls, only plasma {beta}-cryptoxanthin was positively associated with plasma progesterone concentrations (p=0.02). Our findings of heterogeneous associations between tocopherols, carotenes, retinol and IVF outcomes across the stages of IVF suggest a beneficial effect limited to early outcomes and support a complex and context-dependent role of these compounds in female reproduction. This manuscript has been submitted to PlosOne on August 19, 2026.
Liu, H.; Mizani, M. A.; Zhao, Y.; Wood, A.; Inouye, M.; Price, A. L.; Jiang, X.; CVD-COVID-UK/COVID-IMPACT Consortium,
Show abstract
Predicting disease risk from prior diagnoses is fundamental to clinical decision-making, particularly during health emergencies such as the COVID-19 pandemic, when individuals with long-term conditions may be disproportionately vulnerable to adverse outcomes. Despite intense interest in developing models to predict disease risk from prior diagnoses (1-3), most prediction models do not estimate effects of each prior diagnosis on disease risk conditional on other diagnoses, limiting interpretability and clinical utility. We developed the Comorbidity Risk Score (CRS), trained on 13 million individuals (age 40-69) from linked electronic health record (EHR) datasets of the entire population of England, to predict COVID-19 hospitalisation and 87 other disease outcomes. CRS was trained at close to saturated sample size and precisely estimated the effects of 212 prior diagnoses on the 88 disease outcomes, conditional on all other prior diagnoses. Correlations of CRS effect sizes across outcomes (e.g. 0.76 for myocardial infarction vs. hyperlipidaemia) matched the corresponding genetic correlations (e.g. 0.79 for myocardial infarction vs. hyperlipidaemia), confirming that comorbidity architectures capture disease aetiology. On average, CRS identified 5% of the population with 3.4-fold higher disease risk, including myocardial infarction (4.4-fold), lung cancer (6.5-fold), and COVID-19 hospitalisation (6.3-fold). Using prior diagnoses alone, CRS outperformed state-of-the-art clinical COVID-19 models (4). Furthermore, CRS (N=13 million) substantially outperformed state-of-the-art AI (1) (N=0.5 million) and linear (3) (N=0.5 million) models in predicting disease risk, suggesting that training sample size outweighs model complexity. CRS attained near-perfect transferability across self-reported ethnicities (e.g., Black vs. White: AUROC ratio = 97.3%). Finally, CRS distinguished independently predictive comorbidities from indirect associations, e.g., lipid metabolism disorder was a strong predictor of myocardial infarction risk but not ischaemic stroke, after conditioning on other prior diagnoses. In conclusion, CRS provides a comprehensive resource for understanding the impact of comorbidities on COVID-19 and other future diseases, revealing disease aetiology while enabling powerful prediction of disease risk.
Frade, S.; Tunyiswa, Z.; Shin, M.; Dirks, R.
Show abstract
Background: Pressure ulcers often develop complex three-dimensional morphologies that extend beyond the visible wound surface. Subsurface extensions such as tunneling and undermining create hidden cavities that complicate clinical assessment and wound management. Despite their clinical relevance, the prevalence and spatial characteristics of these subsurface wound morphologies have not been well characterized at scale. Methods: We performed a registry-based analysis using data from the LIFT-OFF Pressure Ulcer Registry, which captures longitudinal clinical documentation of pressure ulcers treated in routine care. The registry included approximately 18,000 patients with 32,000 documented pressure ulcers. Spatial characteristics of tunneling and undermining were analyzed using measurements recorded during routine wound assessments, including tract length, direction, and circumferential extent. Directional and circumferential distributions of subsurface defects were examined to characterize wound geometry. Results: Tunneling was present in 764 of 14,700 full-thickness pressure ulcers (5.2%), whereas undermining occurred in 2,293 wounds (15.6%). Tunneling tracts were typically short and exhibited directional clustering relative to the wound bed. In contrast, undermining demonstrated broader circumferential distributions and frequently involved larger subsurface separations beneath the wound margin. Both morphologies demonstrated distinct spatial patterns across anatomical locations and wound stages. Conclusion: Tunneling and undermining are common subsurface features of pressure ulcers and exhibit distinct spatial geometries. Whereas tunneling manifests as directional tract-like extensions, undermining more frequently produces circumferential tissue separation beneath wound margins. Improved characterization of subsurface wound architecture may enhance assessment of wound complexity and provide information not captured by surface measurements alone. Future studies should evaluate whether these features contribute to wound severity assessment, prognosis, and risk stratification.
Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.
Show abstract
Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.